Papers with systematically testing VLMs
UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models (2024.emnlp-main)
Copied to clipboard
Xinyu Pi, Mingyuan Wu, Jize Jiang, Haozhen Zheng, Beitong Tian, ChengXiang Zhai, Klara Nahrstedt, Zhiting Hu
| Challenge: | Vision-Language Models (VLMs) perform on par with larger models in general domain visual grounding and question-answering benchmarks. |
| Approach: | They propose a "Uncontextualized Uncommon Objects" benchmark to evaluate their performance on common datasets. |
| Outcome: | The proposed benchmark focuses on systematically testing VLMs with both large and small parameter counts on rare and specialized objects. |